Papers with memory consumption

43 papers
Teach the Rules, Provide the Facts: Targeted Relational-knowledge Enhancement for Textual Inference (2021.starsem-1)

Copied to clipboard

Challenge: InferBERT is a method to enhance transformer-based inference models with relevant relational knowledge.
Approach: They propose to enhance transformer-based inference models with relevant relational knowledge by injecting relevant facts at test time into the model.
Outcome: The proposed method outperforms existing models on the challenge datasets while outperforming existing models.
Length-Aware Multi-Kernel Transformer for Long Document Classification (2024.starsem-1)

Copied to clipboard

Challenge: Existing SOTA models segment long texts into equal-length snippets, but they have new challenges of context fragmentation and generalizability due to sentence boundaries and varying text lengths.
Approach: They propose a Length-Aware Multi-Kernel Transformer to encode long documents by transformers and vectorize text length by the kernels to promote model robustness over varying document lengths.
Outcome: The proposed model outperforms existing models on five benchmarks from health and law domains up to an absolute 10.9% improvement.
Trove: A Flexible Toolkit for Dense Retrieval (2026.eacl-demo)

Copied to clipboard

Challenge: Existing retrieval tools require considerable engineering effort for many tasks like efficient data management or model customization.
Approach: They propose a novel open-source retrieval toolkit that simplifies research experiments without sacrificing flexibility or speed.
Outcome: The proposed tool reduces memory consumption by 2.6 and allows for arbitrary customizations.
You Only Need Attention to Traverse Trees (P19-1)

Copied to clipboard

Challenge: Recent research has focused on sentence representations.
Approach: They propose a tree-based model that captures phrase-level syntax and word-level dependencies by doing recursive traversal with attention.
Outcome: a new model captures phrase-level syntax and word-level dependencies with attention.
BLINK with Elasticsearch for Efficient Entity Linking in Business Conversations (2022.naacl-industry)

Copied to clipboard

Challenge: Existing systems that align textual mentions of entities to knowledge bases are difficult to deploy in production environments.
Approach: They propose a neural entity linking system that connects entities in business phone conversations to their corresponding Wikipedia and Wikidata entries.
Outcome: The proposed system improves inference speed and memory consumption while maintaining high accuracy.
EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices (2025.acl-industry)

Copied to clipboard

Challenge: Existing KV cache optimizations struggle with irreversible token eviction in long-output tasks . alternative sequence modeling architectures prove costly to adopt within established Transformer infrastructures.
Approach: They propose a memory-efficient solution for infinite contexts that integrates compressed memory into Transformer-based LLMs through a trainable memory-gating module.
Outcome: The proposed solution achieves comparable performance to baseline Transformer-based LLMs while optimizing memory consumption and time to first token.
KV-Latent: Dimensional-level KV Cache Reduction with Frequency-aware Rotary Positional Embedding (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) based on Transformer Decoders have become the preferred choice for conversational generative AI.
Approach: They propose a paradigm called KV-Latent to reduce the KV cache footprint and improve inference speed by down-sampling the Key-Value vector dimensions into a latent space.
Outcome: The proposed paradigm reduces the KV Cache footprint and improves inference speed with a small amount of extra training, less than 1% of pre-training takes.
Parameter-Efficient Fine-Tuning via Circular Convolution (2025.findings-acl)

Copied to clipboard

Challenge: Low-Rank Adaptation (LoRA) has gained popularity for fine-tuning large foundation models, but its intrinsic low-rank characteristic may limit its performance.
Approach: They propose a low-rank adaptive method that uses low-ranked matrices to represent weight changes.
Outcome: The proposed method reduces trainable parameters and mitigates heavy memory consumption associated with full delta matrices by sequentially multiplying mathbf A and mathbb B with the activation.
DREEAM: Guiding Attention with Evidence for Improving Document-Level Relation Extraction (2023.eacl-main)

Copied to clipboard

Challenge: Document-level relation extraction (DocRE) is a task of identifying relations between entities in a document. evidence retrieval (ER) in DocRE faces two major issues: high memory consumption and limited availability of annotations.
Approach: They propose a memory-efficient approach that uses evidence as the supervisory signal . they propose er self-training to learn ER from automatically-generated evidence .
Outcome: The proposed method exhibits state-of-the-art performance on the DocRED benchmark . it uses evidence as the supervisory signal and self-trains on massive data without annotations .
Binarized LSTM Language Model (N18-1)

Copied to clipboard

Challenge: Long short-term memory (LSTM) language models are widely used for automatic speech recognition and natural language processing (NLP) however, they are limited by the word embedding layer.
Approach: They propose to encode words into binary vectors and use binarized LSTM parameters to achieve high memory compression.
Outcome: The proposed model achieves 11.3 compression ratio without loss of performance and 31.6 compression ratio with acceptable performance degradation.
PyramidInfer: Pyramid KV Cache Compression for High-throughput LLM Inference (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods to reduce memory usage for large language models neglect inter-layer dependency between layers and huge memory consumption in pre-computation.
Approach: They propose a method that compresses the KV cache by layer-wise retaining crucial context.
Outcome: The proposed method reduces memory usage by layer-wise retaining crucial context . it can improve 2.2x throughput compared to Accelerate with over 54% memory reduction .
STAR: Constraint LoRA with Dynamic Active Learning for Data-Efficient Fine-Tuning of Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies show that supervised training is still necessary for complex reasoning tasks.
Approach: They propose a method to integrate uncertainty-based active learning and LoRA to effectively integrate the two methods.
Outcome: The proposed approach outperforms baseline models on three reasoning tasks.
Sharper and Faster mean Better: Towards More Efficient Vision-Language Model for Hour-scale Long Video Understanding (2025.acl-long)

Copied to clipboard

Challenge: Existing multimodal large language models (LLMs) have shown impressive performance on the video understanding task, but extremely long videos still pose significant challenges to their context length, memory consumption, and computational complexity.
Approach: They propose a vision-language model named Sophia for long video understanding which can efficiently handle hour-scale long videos.
Outcome: The proposed model exhibits competitive performance compared to existing video understanding baselines across various benchmarks for long video understanding with reduced time and memory consumption.
Blockwise Self-Attention for Long Document Understanding (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in pre-training and fine-tuning methods have drastically reshaped the landscape of natural language processing research.
Approach: They propose a lightweight BERT model that introduces sparse block structures into the attention matrix to reduce memory consumption and training/inference time.
Outcome: The proposed model uses 18.7-36.1% less memory and 12.0-25.1% more time to learn compared to an advanced BERT-based model, RoBERTa.
QSpec: Speculative Decoding with Complementary Quantization Schemes (2025.emnlp-main)

Copied to clipboard

Challenge: Quantization is widely adopted to accelerate inference and reduce memory consumption in large language models.
Approach: They propose a quantization paradigm that decouples efficiency from quality by integrating two complementary schemes via speculative decoding.
Outcome: The proposed approach achieves 1.64x speedup without quality degradation and outperforms state-of-the-art speculative decoding methods by 1.55x in batched settings.
QPruner: Probabilistic Decision Quantization for Structured Pruning in Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Structured pruning can reduce model size but results in significant accuracy degradation . quantization and pruning increase the difficulty of fine-tuning, requiring a more refined quantization scheme.
Approach: They propose a structured pruning framework followed by a layer-wise mixed-precision quantization scheme to reduce model memory consumption during fine-tuning and inference.
Outcome: Experiments on benchmark datasets show that QPruner outperforms existing methods in memory savings while maintaining or improving model performance.
QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language Models (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are quantized to lower precision to reduce memory cost and latency in inference.
Approach: They propose a quantized zeroth-order framework for fine-tuning Large Language Models (LLMs) using low-precision forward passes.
Outcome: The proposed method achieves comparable results to first-order methods in FP8 and superior accuracy in INT8 and INT4 training.
Personalized Federated Learning for Text Classification with Gradient-Free Prompt Tuning (2024.findings-naacl)

Copied to clipboard

Challenge: Pretrained language models (PLMs) are used for personalized federated learning . communication costs are high with large PLMs, and local training is expensive .
Approach: They propose a framework for federated learning with pretrained language models . they propose 'discrete local search' and compression mechanism for local training .
Outcome: The proposed framework achieves superior performance compared with baselines.
Text Level Graph Neural Network for Text Classification (D19-1)

Copied to clipboard

Challenge: Recent researches have explored graph neural network (GNN) techniques on text classification, but they are faced with the problems of fixed corpus level graph structure which don’t support online testing and high memory consumption.
Approach: They propose a graph neural network model that builds graphs for each input text with global parameters sharing instead of a single graph for the whole corpus.
Outcome: The proposed model outperforms existing models on several text classification datasets even with consuming less memory.
SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget (2024.acl-long)

Copied to clipboard

Challenge: Mixture of experts (MoE) is a popular technique to improve capacity of Large Language Models (LLMs) but memory-constrained devices are a major concern in edge AI training and serving.
Approach: They propose a framework for efficient serving of MoE-based large language models with tunable memory budgets.
Outcome: Experiments show that SwapMoE can reduce memory consumption while maintaining reasonable accuracy.
Phrase-level Self-Attention Networks for Universal Sentence Encoding (D18-1)

Copied to clipboard

Challenge: Phrase-level self-attention networks (PSAN) can capture context dependencies at the phrase level instead of the sentence level.
Approach: They propose to perform self-attention across words inside a phrase to capture context dependencies at the phrase level and use the gated memory updating mechanism to refine each word’s representation hierarchically with longer-term context dependency captured in a larger phrase.
Outcome: The proposed model can achieve state-of-the-art performance across a plethora of NLP tasks including binary and multi-class classification, natural language inference and sentence similarity.
DynamicKV: Task-Aware Adaptive KV Cache Compression for Long Context LLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing KV cache compression methods enforce a fixed pattern, neglecting task-specific characteristics, which hampers the effective retention of essential information while discarding less important tokens.
Approach: They propose a Task-Aware KV cache mechanism that dynamically adjusts the KV caching size across different layers based on the characteristics of the tasks.
Outcome: The proposed method surpasses state-of-the-art methods by 11% on the LongBench dataset even under extreme compression (0.9%)
NACL: A General and Effective KV Cache Eviction Framework for LLM at Inference Time (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) with extended context windows are expensive and infeasible on fixed memory hardware due to the surprisingly large memory consumption of KV Cache.
Approach: They propose a general framework for long-context KV cache eviction that achieves more optimal and efficient evict in a single operation during the encoding phase.
Outcome: The proposed framework improves performance on short- and long-text tasks by 80% and 76% respectively, reducing KV Cache by up to 5 with over 95% performance maintenance.
HiTIN: Hierarchy-aware Tree Isomorphism Network for Hierarchical Text Classification (2023.acl-long)

Copied to clipboard

Challenge: Existing dual-encoder methods in HTC achieve weak performance gains with huge memory overheads and their structure encoders heavily rely on domain knowledge.
Approach: They propose a hierarchy-aware tree isomorphism network to enhance the text representations with only syntactic information of the label hierarchy.
Outcome: The proposed model could boost the performance of hierarchical text classification without prior statistics or label semantics without prior data.
XQuant: Achieving Ultra-Low Bit KV Cache Quantization with Cross-Layer Compression (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks. however, their extensive memory requirements present significant challenges for deployment in resource-constrained environments.
Approach: They propose a training-free framework that achieves ultra-low equivalent bit-width KV cache quantization.
Outcome: The proposed framework outperforms state-of-the-art methods on TruthfulQA and LongBench.
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Quantization-aware training (QAT) is a low-bit training solution that requires substantial training resources.
Approach: They propose an algorithm that reduces memory consumption by low-bit representations with minimal accuracy loss.
Outcome: EfficientQAT achieves 2-bit Llama-2-70B model on single GPU in 41 hours . compared to previous methods, it obtains model with less than 3 points accuracy degradation .
SANCL: Multimodal Review Helpfulness Prediction with Selective Attention and Natural Contrastive Learning (2022.coling-1)

Copied to clipboard

Challenge: e-commerce has become a research hotspot for review helpfulness prediction . a new approach to help predict helpfulness of multimodal product reviews is proposed .
Approach: They propose a machine learning task to identify helpfulness of multimodal product reviews . they use a probe-based strategy to enforce high attention weights on regions of greater significance .
Outcome: The proposed model achieves state-of-the-art performance with lower memory consumption on two benchmark datasets with three categories.
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts (2025.acl-long)

Copied to clipboard

Challenge: Existing approaches to optimize large language models for long-context inference are inefficient and consume memory.
Approach: They propose a mixed-precision quantization method via mixture of experts that inputs tokens into router chunk by chunk to reduce inference overhead.
Outcome: The proposed method outperforms state-of-the-art KV cache quantization methods on multiple benchmark datasets.
LoRAExit: Empowering Dynamic Modulation of LLMs in Resource-limited Settings using Low-rank Adapters (2024.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have exhibited remarkable performance across various natural language processing tasks, but deployment on resource-limited settings remains a challenge.
Approach: They propose a dynamic inference architecture that leverages low-rank adaptors for efficient deployment of LLMs.
Outcome: The proposed architecture significantly improves performance when deployed on resource-limited settings.
Q-Mamba: Towards more efficient Mamba models via post-training quantization (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies show that Mamba architectures have room for further optimization in linear projections and state caches.
Approach: They propose a decoupled scale quantization scheme to mitigate outliers in states and channels by applying separate quantization scales.
Outcome: The proposed method reduces memory consumption by 50% across various quantization settings, model sizes, and generation and zero-shot tasks.
NeuroAda: Activating Each Neuron’s Potential for Parameter-Efficient Fine-Tuning (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods for parameter-efficient fine-tuning are limited and require computational and memory resources.
Approach: They propose a parameter-efficient fine-tuning method that enables fine-grained model finetunation while maintaining high memory efficiency.
Outcome: The proposed method reduces CUDA memory usage by up to 60% while maintaining high performance.
ResFormer: All-Time Reservoir Memory for Long Sequence Classification (2025.emnlp-main)

Copied to clipboard

Challenge: Existing models with quadratic time and memory complexity restrict input length . however, analyzing extensive sequential contexts is challenging .
Approach: They propose a neural network architecture that captures contextual dependencies in linear time and a nonlinear readout to model short-term dependencies within sentences.
Outcome: The proposed model outperforms baseline models on EmoryNLP datasets and on IEMOCAP and MultiWOZ datasets.
Layer-Condensed KV Cache for Efficient Inference of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Using a key-value cache, memory consumption is a bottleneck for high-throughput language models.
Approach: They propose a method that only computes and caches the KVs of a small number of layers, thus saving memory consumption and improving inference throughput.
Outcome: The proposed method achieves higher throughput and competitive performance than standard transformers and is orthogonal to existing transformer memory-saving techniques.
A Comprehensive Evaluation of Quantization Strategies for Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Quantization studies have focused on instruction-tuned LLMs, leaving their performance on other benchmarks unclear.
Approach: They propose a framework to evaluate quantized large language models using four dimensions . they propose to reduce the bits needed for model weights or activations with minimal performance loss .
Outcome: The proposed framework can retain comparable performance to non-quantized LLMs on most benchmarks.
Simplified Graph Learning for Inductive Short Text Classification (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for short text classification are limited and lack of labeled data is not enough.
Approach: They propose a novel short text classification algorithm which leverages words to handle the lack of labeled data.
Outcome: The proposed model performs better with lower memory consumption and faster inference speed than previous models.
A Survey for Efficient Open Domain Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Open domain question answering (ODQA) is a longstanding task that can answer factoid questions without explicit evidence in natural language processing (NLP).
Approach: They propose to use open domain question answering to answer factual questions from a large knowledge corpus without explicit evidence.
Outcome: The proposed models can answer factoid questions from a large knowledge corpus without explicit evidence.
QEFT: Quantization for Efficient Fine-Tuning of LLMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to optimize inference and fine-tuning for large language models have failed to improve all aspects of the process.
Approach: They propose a new technique that accelerates both inference and fine-tuning while using fewer resources.
Outcome: The proposed technique accelerates both inference and fine-tuning while using fewer resources.
Reconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting Transcripts (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to meeting summarization are limited due to noise, lengthy transcripts, and scattered salient information.
Approach: They propose a two-step framework for meeting summarization that leverages a self-supervised paradigm to reconstruct transcripts and a relative positional bucketing algorithm to equip models to generate the summary.
Outcome: The proposed method significantly reduces memory consumption and processing time on two meeting summarization datasets.
RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for evicting KV pairs rely on the "persistence of importance" hypothesis . visual tokens display "deferred importance" but become pivotal during later decoding, authors say .
Approach: They propose an entropy-driven method that reformulates KV eviction from "discrete context truncation" to "continuous memory evolution" they propose to prune visual tokens with "deferred importance" visual token exhibiting low salience but becoming pivotal during later decoding .
Outcome: The proposed method achieves 5.0 KV cache compression and 1.5 decoding acceleration.
METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding (2025.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Video Large Language Models (VLLMs) have significantly enhanced their ability to understand video content.
Approach: They propose a training-free, Multi-stage Event-based Token compression framework that eliminates redundant visual tokens across three critical stages .
Outcome: The proposed framework reduces FLOPs and KV Cache memory consumption while maintaining comparable or even superior accuracy.
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture (2025.findings-emnlp)

Copied to clipboard

Challenge: Long-context Large Language Models (MLLMs) are critical for video understanding and image analysis.
Approach: They propose a hybrid architecture that integrates Mamba and Transformer blocks . they introduce data construction methods that capture both temporal and spatial dependencies .
Outcome: The proposed model achieves competitive results across various benchmarks while maintaining high throughput and low memory consumption.
Reducing Peak Memory Usage for Modern Multimodal Large Language Model Pipelines (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to reduce memory usage of multimodal large language models rely on storing large numbers of vision tokens in the key–value cache . however, such compression is typically only applied after all inputs are processed, resulting in high peak memory usage during the prefill stage.
Approach: They propose a sequential input-compression mechanism that enforces a fixed memory budget by performing structure-aware key–value cache compression during the prefill stage.
Outcome: The proposed approach reduces peak memory usage while maintaining generative performance with only minimal degradation, enabling more practical and memory-efficient multimodal inference.
ReFreeKV: Towards Threshold-Free KV Cache Compression (2026.findings-acl)

Copied to clipboard

Challenge: Towards the KV cache efficiency, we propose a new objective that lifts the threshold constraints for robust KV compression.
Approach: They propose a method that adjusts KV cache budgets while preserving full-cache performance.
Outcome: The proposed method can reduce memory consumption while preserving full-cache performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations